The real goal is never accuracy on training data
When you train a machine learning model, you optimise its performance on the data you showed it. But that is not actually what you care about. You care about what the model does on data it has never seen. Will it work on tomorrow's emails? On next month's customer transactions? On medical images from a different hospital?
That ability to perform well on new, unseen data is called generalisation. It is the actual goal. Training accuracy is just a proxy metric, and a misleading one if you are not careful.
Imagine a student preparing for an exam by memorising the exact questions from last year's past papers, word for word. They ace the practice papers. But the real exam has different questions covering the same material. The memoriser fails while the student who actually understood the concepts does fine. Overfitting is the machine learning equivalent of that kind of memorising.
This is why you always split your data into a training set and a test set before you start. The test set is data the model never sees during training. It is your honest proxy for how the model will perform in the real world. You train on the training set. You evaluate on the test set. You never use the test set to make decisions about your model. The moment you do, it is no longer an honest measure of generalisation.
Underfitting and overfitting
Most problems in machine learning sit on a spectrum between two failure modes. On one end, the model is too simple. On the other, it is too complex.
Model is too simple to capture the real pattern. It misses obvious structure in the training data and performs poorly on both training and test data.
Model captures the true underlying pattern without getting distracted by noise. Works well on training data and generalises to new data.
Model is too complex and follows the training data perfectly, including its noise and random quirks. Fails badly on new data it hasn't seen before.
Underfitting is usually obvious and easy to fix: your model is too simple, so you make it more complex. Add more features, use a deeper tree, try a more powerful algorithm. The challenge is that fixing underfitting tends to move you toward overfitting, and vice versa. Finding the right point in between is what model development is really about.
The bias-variance tradeoff
Every model has two sources of error: bias and variance. These two tend to pull in opposite directions, and understanding them helps you diagnose problems and choose solutions.
Bias is systematic error from wrong assumptions. A linear model trying to fit a curved pattern will always be off, no matter how much data you give it. It is biased toward straight lines. More data won't fix this; you need a more flexible model.
Variance is sensitivity to the specific training data. A highly flexible model that follows every wiggle in the training data will produce very different results if you retrain it on a slightly different sample. It is unstable. It has memorised quirks that don't generalise.
As model complexity increases, bias (blue) decreases but variance (red) increases. Total error (black) is minimised at the sweet spot where neither dominates. The goal is to find a model complex enough to capture the pattern, but not so complex that it memorises the noise.
How to fix overfitting
Overfitting is by far the more common problem in practice, because modern models are powerful enough to memorise any finite dataset if you let them. Here are the main weapons against it.
Cross-validation: a more honest evaluation
When your dataset is small, a single train-test split can be noisy. If you happen to put all the easy examples in the test set, your accuracy looks great but means nothing. Cross-validation solves this by rotating which data gets used for testing.
In k-fold cross-validation, you split the data into k equal parts (folds). You train on k-1 folds and test on the remaining one. You repeat this k times, each time using a different fold as the test set. You average the results across all k rounds. This gives you a much more reliable estimate of generalisation performance with no data wasted.
from sklearn.datasets import load_iris from sklearn.tree import DecisionTreeClassifier from sklearn.model_selection import cross_val_score import numpy as np X, y = load_iris(return_X_y=True) # Compare three models: shallow, medium, and deep trees for depth in [1, 4, 20]: model = DecisionTreeClassifier(max_depth=depth, random_state=42) scores = cross_val_score(model, X, y, cv=5) print( f"Depth {depth:<3} | " f"CV accuracy: {scores.mean():.3f} (+/- {scores.std():.3f})" )
Depth 4 | CV accuracy: 0.953 (+/- 0.022) ← good fit
Depth 20 | CV accuracy: 0.940 (+/- 0.038) ← slightly overfitting
Notice that depth 20 is actually slightly worse than depth 4 on average, and more variable (higher standard deviation). The model at depth 20 is overfitting. Depth 4 finds the sweet spot for this dataset. Cross-validation makes this visible in a way that a single train-test split often cannot.
Never touch your test set until you have finished making all decisions about your model. Every time you check test performance and then change something, the test set leaks information into your development process. It is no longer an honest measure of generalisation. If you need to tune hyperparameters, use a separate validation set or cross-validation on the training data only. The test set is a one-shot, final-answer measure.
"The most important skill in machine learning is resisting the urge to keep checking your test score."
A principle learned repeatedly, the hard way